BMC Medical Genomics
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match BMC Medical Genomics's content profile, based on 50 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Goroshchuk, O.; Koller, D.
Show abstract
Background: Endometriosis affects approximately 10% of reproductive-age women and is associated with substantial diagnostic delay and heterogeneous symptom presentation. Prior machine-learning prediction models have relied on comorbidity data alone or on small candidate-variant genetic scores, with inconsistent or incompletely reported performance. No study has combined a well-powered, multi-ancestry polygenic risk score (PRS) with environmental, reproductive, and symptom data in a single hybrid model. We developed and evaluated hybrid risk-prediction models integrating a genome-wide, multi-ancestry PRS with clinical and symptom data for endometriosis in the US-based All of Us Research Program. Methods: Among 69,376 participants (15,382 endometriosis cases, 53,994 controls) across six genetically inferred ancestry groups, we computed individual-level PRS values using PRS-CS weights derived from an independent, multi-ancestry GWAS. Five nested logistic regression, random forest, and XGBoost models progressively added age, ancestry, and within-ancestry genetic principal components (Model 1), environmental and reproductive factors (Model 2), symptom and comorbidity indicators (Model 3), all covariates combined (Model 4), and PRS x environment interactions (Model 5). Performance was assessed by AUROC in a held-out test set and 5-fold cross-validation, with class-weighted, Youden-optimized thresholds used for sensitivity, specificity, and predictive values; permutation importance identified top contributors. Pairwise AUROC differences were tested with a Holm-corrected DeLong-type test. Results: Discrimination improved from AUROC 0.63 (PRS, age, ancestry, principal components) to 0.72 for the full model, driven mainly by symptom and comorbidity data. XGBoost consistently outperformed logistic regression and random forest. The PRS ranked among the top individual predictors by permutation importance in nearly every model, alongside age, while genetic and demographic information alone gave only modest discrimination, and PRS x environment interactions did not improve on environmental factors alone. Threshold optimization yielded balanced sensitivity and specificity (~0.67/0.65) versus near-zero sensitivity at a default threshold. Conclusions: Combining the PRS with symptom and comorbidity data gave the best discrimination compared to solely a well-powered, multi-ancestry PRS as a predictor of endometriosis. This study clarifies both the promise and current limits of hybrid genetic-clinical prediction for endometriosis and points to symptom-based phenotyping, molecular subtyping, and external validation as priorities.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Chia, C.; Baker, K.
Show abstract
Obesity is a significant public health concern. Early-onset obesity in the context of rare disease can reflect genetically-mediated pathology or elevated susceptibility through indirect mechanisms. Mapping the diverse characteristics and needs of young people with obesity in the rare disease population is a first step toward mechanistic and translational research. We carried out a retrospective comparative analysis of demographic, genotypic, phenotypic and health service utilisation data for young people with obesity (cases: n=500) and without obesity (controls: n=11,444) from the UK 100,000 Genomes Project rare disease cohort. Cases and controls were recruited prior to genomic diagnosis, across clinical disorder categories. We observed significant association between socioeconomic deprivation and obesity risk. Young people with obesity had significantly higher utilisations of acute care and mental health services, indicating an overall higher health burden. A curated panel of 519 candidate obesity-associated genes demonstrated aggregate association with obesity, although no single gene reached significance. Phenotypic comparison between cases and controls highlighted increased multi-organ and neurological system involvement, highlighting the overlap between neurodevelopmental and obesity risks. Within the case group, we conducted cluster analysis to identify early-onset obesity groups with different phenotypic profiles, potentially arising from different causal pathways - this identified six obesity subgroups of interest, with differing involvement of neurodevelopmental and other systems. Our study confirms that obesity co-occurs with a wide range of factors within the rare disease population, and is associated with significant physical and mental health needs, requiring holistic lifelong care.
Lin, N.; Balasubramanian, R.; Menichetti, G.; Eliassen, H.; Trabert, B.; Avila-Pacheco, J.; Townsend, M. K.; Terry, K. L.; Clish, C. B.; Tworoger, S. S.; Zeleznik, O. A.
Show abstract
Background: Evidence suggests chronic distress influences ovarian cancer (OC) etiology and metabolomic profiles. Here, we evaluated the association of a metabolite-based distress score (MDS) and OC risk. Methods: We included two matched case-control studies nested within the Nurses' Health Studies (N=584) and the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial (N=348). Metabolites were measured 3-27 years before diagnosis using liquid-chromatography tandem mass spectrometry. We examined the association of quintiles of MDS and 19 constituent metabolites with OC risk using unconditional logistic regression and stratified by tumor histotype, menopausal status, and age at diagnosis. Results: We observed women in the highest versus lowest quintile of MDS had an increased OC risk (OR=1.62,95%CI=1.03-2.54,ptrend=0.07), and type 2 tumors (OR=1.71,95%CI=1.03-2.83,ptrend=0.11). Associations were suggestively stronger for premenopausal and <69-year-old women, and driven by pseudouridine, and N2,N2-dimethylguanosine. Conclusion: Our findings suggest chronic distress-associated metabolic dysregulation may represent a novel OC risk factor, especially among younger women.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Venkatesh, R.; Deo, R.; Cappola, T.; Penn Medicine BioBank, ; Ritchie, M. D.; Kim, D.
Show abstract
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and a major cause of cardioembolic stroke. Although polygenic risk scores (PRS) are well characterized to quantify inherited susceptibility for AF, they provide limited insight into the pathways and tissues underlying genetic risk, which are critical to uncover for individual risk prediction. In this study, we develop a pathway-level multi-omics representation learning framework that converts individual genetic profiles into interpretable biological features by integrating GWAS-derived pathway burden scores with tissue-specific transcriptomic pathway signals. We constructed machine learning models to assess population-level AF risk prediction performance across genomic and transcriptomic tissue contexts; the pathway-based global attention models substantially improved risk prediction performance over PRS and other baselines (AUROC improved from 0.601 to 0.738). Transformer and graph neural network frameworks then assessed individual-level pathway interpretability, revealing heterogeneous contributions from electrical signaling, cardiac development, and DNA repair pathways to AF risk. This added interpretability highlights the potential of this pathway approach to enable more mechanistically informed risk stratification than static PRS by capturing underlying heterogeneity. To independently assess whether prioritized pathways reflected cardiac regulatory biology, we compared pathway rankings with transcriptional effects predicted by the AlphaGenome foundation model. Variants in highly ranked pathways showed significantly greater predicted effects on expression in atrial and ventricular tissues (FDR = 0.032) relative to controls, providing orthogonal evidence that the model identifies biologically relevant mechanisms. Overall, this work reframes polygenic risk from a single measure of susceptibility to tissue-informed pathway mechanisms, providing a framework for interpretable genomic stratification in complex diseases.
Shachar, E. K.; Haas, R.; Rodriguez, V. E.; Lester, J.; Siavoshi, M. A.; Kwan, L.; Niell-Swiller, M.; Spellman, P. T.; Boutros, P. C.; Chang, V. Y.; Karlan, B. Y.
Show abstract
Importance: Chronic stress may contribute to adverse health outcomes through cumulative physiologic dysregulation. Allostatic load (AL), a composite measure of multisystem physiologic burden, may capture biologic effects of structural, social, and psychosocial stress not reflected by self-reported measures. Objective: To evaluate racial and ethnic differences in AL among women with familial cancer risk and examine how socioeconomic status, psychosocial factors, clinical characteristics, and health behaviors contribute to variations in AL. Design: Cross-sectional study of underrepresented minority participants enrolled in the HERSTORY cohort from October 2023 through September 2025, with comparison participants from the UCLA ATLAS biobank. Setting: UCLA academic health system. Participants: The study included 303 racially and ethnically diverse female HERSTORY participants aged [≥]35 years with a family history of cancer and matched non-Hispanic White female ATLAS participants (n=709). Exposures: Race and ethnicity, age, neighborhood deprivation, cancer history and stage, depression, perceived stress, cancer worry, and physical activity. Main Outcomes and Measures: The primary outcome was AL, calculated from cardiometabolic and organ-function measures. A secondary index incorporated race- and ethnicity-specific neutrophil-to-lymphocyte ratio (NLR) derived from 326,826 women in the UCLA Health population. Multivariable regression models evaluated factors associated with elevated AL. Results: Compared with matched non-Hispanic White participants, Black and Asian/Pacific Islander HERSTORY participants had significantly higher AL after adjustment. Hispanic/Latina participants did not have significantly elevated AL. Older age, greater area-level socioeconomic deprivation, and depression were independently associated with higher AL. Prior cancer diagnosis, cancer worry and perceived stress were not significantly associated with AL, whereas regular physical activity was associated with lower AL. Among cancer patients, advanced stage was associated with greater AL. Conclusions and Relevance: This study demonstrates elevated AL among understudied racial/ethnic minority groups with familial cancer risk and identifies associations with neighborhood deprivation, depression, and physical activity. The association between cancer stage and AL suggests that physiologic stress may reflect variation in cancer burden. The lack of association with perceived stress and cancer worry further indicates that physiologic and self-reported psychosocial measures capture distinct dimensions of stress. The development of race/ethnicity-specific NLR thresholds derived from large population samples provide a benchmark for future studies.
Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.
Show abstract
Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.
Show abstract
Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.
Show abstract
Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.
Mathews, R.; Bouyadjera, S. B.; Donegan, J. J.; Havird, J. C.
Show abstract
Mitochondria are central hubs for cellular metabolism and mitochondrial dysfunction is a hallmark of many chronic diseases. Consequently, changes in mitochondrial DNA copy number (mtDNA-CN), the number of mtDNA genomes per cell or tissue sample, are associated with diseases ranging from cancer and obesity to psoriasis and all-cause mortality. MtDNA-CN especially holds promise as a biomarker for neurodegenerative diseases, but whether and how mtDNA-CN changes with neurodegeneration is controversial. Here, we performed a systematic review and meta-analysis of 76 studies including 156 comparisons of mtDNA-CN in populations with or without a neurodegenerative disease to identify overall trends and potential moderators that explain variation among studies. Overall, mtDNA-CN was not statistically different with neurodegeneration, but heterogeneity among studies was extreme (I2 = 99.5%). The diagnosed disease explained the most variation. For example, Alzheimer's patients showed a 21% decrease in mtDNA-CN, but there was no change in mtDNA-CN with Parkinson's disease. Decreases in mtDNA-CN during neurodegeneration were also more extreme at older ages. Surprisingly, the tissue sampled for mtDNA-CN was not particularly influential, except for certain diseases. Studies published in earlier years also showed more extreme decreases in mtDNA-CN with neurodegeneration. Excessive heterogeneity persisted even after accounting for all moderators and their interactions (I2 = 85.7%). We conclude that the general perception of decreased mtDNA-CN with neurodegeneration is a vast oversimplification that may stem from legacy effects of early studies. However, mtDNA levels offer great promise as biomarkers for neurodegeneration, other diseases, and general health metrics, assuming appropriate complications can be considered.
Bresnahan, S. T.; Xiong, C.; Head, T.; Chang, Y.-H.; Bhattacharya, A.; Huang, J. Y.
Show abstract
Unmeasured confounding threatens causal inference and replicability in observational multi-omic studies across variable environments. Genetic instrumental variables (Mendelian randomization) and negative-control calibration each address complementary sources of unmeasured confounding, yet no existing framework unifies them for omics-scale mediation analysis. We introduce ICONIC, an R package that embeds genetic instruments and negative controls within a proximal causal inference framework for total-effect and mediation analysis. ICONIC implements eight estimators spanning five confounding-control strategies, supports continuous, binary, and time-to-event outcomes, and provides extensive diagnostics including sensitivity analyses that map estimator performance across plausible assumptions. Ground-truth benchmarks are calibrated to real-omics covariance structures via a hybrid generative model (GAN + feature-level Gaussian copula) rather than parametric simulation, and a companion planning tool predicts performance gains from collecting additional omic data. We demonstrate ICONIC in two case studies: identifying placental transcriptomic mediators of gestational diabetes on birth weight (n = 164), and tumor-expression mediators of smoking intensity on lung cancer survival (n = 494). Notably, ICONIC's diagnostics recommended different estimation strategies across the two scenarios, reflecting differences in the likely influence of unmeasured confounding. ICONIC is freely available at https://github.com/sbresnahan/iconic/.
Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.
Show abstract
Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.
Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.
Show abstract
Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.